首页> 外文OA文献 >Bringing ultra-large-scale software repository mining to the masses with Boa

【2h】

Bringing ultra-large-scale software repository mining to the masses with Boa

机译：通过Boa将超大规模软件存储库挖掘带给大众

代理获取

本网站仅为用户提供外文OA文献查询和代理获取服务，本网站没有原文。下单后我们将采用程序或人工为您竭诚获取高质量的原文，但由于OA文献来源多样且变更频繁，仍可能出现获取不到、文献不完整或与标题不符等情况，如果获取不到我们将提供退款服务。请知悉。

页面导航

摘要
著录项
引文网络
相似文献
相关主题

摘要

Mining software repositories provides developers and researchers achance to learn from previous development activities and apply thatknowledge to the future. Ultra-large-scale open source repositories(e.g., SourceForge with 350,000+ projects, GitHub with 250,000+projects, and Google Code with 250,000+ projects) provide an extremelylarge corpus to perform such mining tasks on. This large corpus allowsresearchers the opportunity to test new mining techniques andempirically validate new approaches on real-world data. However, thebarrier to entry is often extremely high. Researchers interested inmining must know a large number of techniques, languages, tools, etc,each of which is often complex. Additionally, performing mining atthe scale proposed above adds additional complexity and often isdifficult to achieve.The Boa language and infrastructure was developed to solve theseproblems. We provide users a domain-specific language tailored forsoftware repository mining and allow them to submit queries via ourweb-based interface. These queries are then automaticallyparallelized and executed on a cluster, analyzing a dataset containingalmost 700,000 projects, history information from millions ofrevisions, millions of Java source files, and billions of AST nodes.The language also provides an easy to comprehend visitor syntax toease writing source code mining queries. The underlyinginfrastructure contains several optimizations, including queryoptimizations to make single queries faster as well as a fusionoptimization to group queries from multiple users into a single query.The latter optimization is important as Boa is intended to be ashared, community resource. Finally, we show the potential benefit ofBoa to the community by reproducing a previously published casestudy and performing a new case study on the adoption of Java languagefeatures.

机译：采矿软件存储库为开发人员和研究人员提供了从先前的开发活动中学习并将其知识应用于未来的机会。超大规模开放源代码存储库（例如，具有35万多个项目的SourceForge，具有25万多个项目的GitHub和具有25万多个项目的Google Code）提供了非常庞大的语料库来执行此类挖掘任务。这种庞大的语料库使研究人员有机会测试新的挖掘技术，并根据经验验证实际数据的新方法。但是，进入壁垒通常非常高。有兴趣的研究人员必须了解大量的技术，语言，工具等，其中每种技术通常都很复杂。此外，按上述建议的规模进行挖掘会增加额外的复杂性，而且通常难以实现。开发了Boa语言和基础结构来解决这些问题。我们为用户提供了一种针对特定领域的语言，专门针对软件存储库挖掘而开发，并允许他们通过基于Web的界面提交查询。然后，这些查询将自动并行化并在集群上执行，分析包含近700,000个项目的数据集，来自数百万次修订的历史信息，数百万个Java源文件以及数十亿个AST节点。挖掘查询。基础基础结构包含多项优化，包括使单个查询更快的查询优化以及将来自多个用户的查询分组为单个查询的融合优化。后一个优化非常重要，因为Boa旨在成为共享的社区资源。最后，我们通过复制以前发布的案例研究并就采用Java语言功能进行新的案例研究，来显示Boa对社区的潜在好处。

著录项

作者
Dyer, Robert;
展开▼
作者单位

展开▼
年度 2013
总页数
原文格式 PDF
正文语种 en
中图分类

相似文献

外文文献
中文文献
专利

1. Boa: Ultra-Large-Scale Software Repository and Source-Code Mining [J] . ROBERT DYER, HOAN ANH NGUYEN, HRIDESH RAJAN, ACM transactions on software engineering and methodology . 2016,第1期

机译：Boa：超大型软件存储库和源代码挖掘
2. Mining software repositories for empirical validation of laws of software evolution for Java projects [J] . Arvinder Kaur, Vidhi Vig International journal of computational systems engineering . 2016,第3期

机译：挖掘软件存储库以对Java项目的软件演化定律进行经验验证
3. MSR4SM: Using topic models to effectively mining software repositories for software maintenance tasks [J] . Sun Xiaobing, Li Bixin, Leung Hareton, Information and software technology . 2015,第Octa期

机译：MSR4SM：使用主题模型来有效地挖掘软件存储库用于软件维护任务
4. Boa: A language and infrastructure for analyzing ultra-large-scale software repositories [C] . Dyer Robert, Nguyen Hoan Anh, Rajan Hridesh, International Conference on Software Engineering . 2013

机译：Boa：一种用于分析超大规模软件存储库的语言和基础架构
5. Bringing ultra-large-scale software repository mining to the masses with Boa. [D] . Dyer, Robert. 2013

机译：通过Boa将超大规模软件存储库挖掘带给大众。
6. The Democratization of Diagnosis: Bringing the Power of Medical Diagnosis to the Masses [O] . Rupan Bose, Leslie A. Saxon 2019

机译：诊断的民主化：将医学诊断的力量带给大众
7. Mining Software Repositories with a Collaborative Heuristic Repository [O] . Hlib Babii, Julian Aron Prenner, Laurin Stricker, 2021

机译：使用协作启发式存储库的挖掘软件存储库

Bringing ultra-large-scale software repository mining to the masses with Boa

摘要

著录项

引文网络

相似文献

相关主题

期刊订阅